Papers with ADPO algorithm
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing toxicity within large language models can negatively impact the user experience, causing performance degradation. |
| Approach: | They propose an adversarial DPO algorithm that improves direct preference optimization (DPO) by incorporating harmful data into the generative model. |
| Outcome: | The proposed training algorithm improves the model’s resilience against harmful conversations while minimizing performance degradation. |